AI Learning Series · Part 25

Beyond the Transformer

The new-architectures landscape: state-space models, linear attention, 1.58-bit networks, byte models, diffusion LMs, and decision models — every serious challenger to the autoregressive-softmax stack, and the specific assumption each one drops.

MoE SOTA Survey
→
KV Cache Types
→
New Architectures
→
Latest News

01 The Big Picture

The transformer is not one idea. It is a stack of assumptions bundled together in 2017 — and each newcomer in this doc is an attack on exactly one of them.

Doc 02 built the baseline: an autoregressive model that generates one token at a time, each token attending to every previous token through softmax attention, with the KV cache as its memory. Doc 21 showed the incumbent evolving inside that frame — MoE sparsity, long context, better caches. This doc is the outside view: architectures that don't tune the transformer but replace parts of its skeleton.

The assumptions under attack, one per family:

Transformer assumptionWho drops itWhat they gain
Attention must be O(s²) and exact over all historyMamba / SSMs, RWKV, Gated DeltaNetO(s) time, fixed-size state
Every layer must be attentionJamba, Zamba, Griffin/Hawk hybridsRecall where needed, speed everywhere else
Weights need 16 bitsBitNet b1.58~9× less memory per weight
Text must be chopped into tokens firstByte Latent TransformerNo tokenizer, no token tax
Output must be emitted left-to-rightDiffusion LMs (LLaDA, Mercury)Many tokens per step (bridges doc 24)
The model must generate text to answerJEPA family, JEV decision modelLatents or typed decisions — no decoding loop at all
🧭
Reading rule for this doc: for each family ask three questions — what is the core mechanism and its math, why might it win, and why might it not? Nearly every challenger wins on compute/bandwidth economics and loses on something attention gave for free: exact recall, mature tooling, or generative quality.

02 What — The Taxonomy Up Front

Here is the whole landscape on one card. Generative quality is a judgment about frontier-class text benchmarks, not a moral ranking.

ArchitectureTime / tokenState sizeGenerative qualityRepresentative models
Softmax attention (baseline)O(s²) train, O(s) per new tokenKV cache — grows O(s)★ referenceGPT, Llama, Claude, Gemini
State-space model (selective)O(s)Fixed — one h per layerNear-parity, weaker recallMamba, Mamba-2, Codestral Mamba
Attention–SSM hybridO(s²) + O(s) mixedGrows only at attention layersParity (best of both)Jamba, Zamba, Griffin/Hawk
Linear attention / DeltaNetO(s)Fixed matrix S (d×d)Good LM, weak multi-query recallGated DeltaNet, Kimi Linear
Data-dependent linear RNNO(s)Fixed per-token stateGood LM, weak recallRWKV v5/v6
Exponential-gate LSTMO(s)Fixed cell + matrix memoryCompetitive at small–mid scalexLSTM (sLSTM/mLSTM)
Ternary-weight LMSame O(s²) — cheaper opsSame KV cacheParity at matched scaleBitNet b1.58, b1.58 2B4T
Byte-level patching LMO(bytes + patches²)KV over patchesParity at large scaleByte Latent Transformer
Diffusion LMO(s²) per refinement passFull-sequence activationsClosing gap; weaker long-formLLaDA, Mercury (Inception), Grok-fast style
Joint-embedding predictiveO(context) — no decodeLatent representationN/A — not generativeI-JEPA, V-JEPA (LeCun lineage)
Decision model (typed readout)O(S) prefill, O(1) per questionShared-context KV, reusedN/A — calibrated decisions, not textJEV (TypeSafe)
Memory-augmented transformerO(s²) + NN lookupWeights + product-key storeParity + long-tail factsProduct-key memory layers, Titan

MoE is deliberately absent: it changes which parameters fire, not the sequence math, and got its full survey in doc 21. The rows above change the sequence math itself.

03 Why These Arise Now

The quadratic wall. Attention computes s² pairwise scores per layer. At s = 128K that is ~16 billion score computations per head — training cost grows with the square, and doc 22 showed the KV cache eating GPU memory linearly with context. Linear-recurrence families collapse this to O(s).
Bandwidth economics. Doc 10 established that decoding is memory-bandwidth-bound: each token reads the whole model + KV cache from HBM. Anything that shrinks state (BitNet's 1.58-bit weights, SSM's fixed-size h) attacks the real bottleneck — bytes moved, not FLOPs.
Extrapolation limits. Softmax attention trained at length L degrades past L. Recurrent state is length-agnostic by construction — the same h update runs at position 10 or 10 million.
The decoding tax. Autoregressive decoding emits one token per full forward pass (doc 24). Diffusion LMs and multi-token heads attack the sequentiality itself; decision models skip the loop entirely when the answer is a choice, not prose.
The tokenizer tax. BPE fragments code, math, and non-English text; token count ≠ character count. Byte-level models delete the tokenizer and pay compute instead.

04 How — Computing a 4-Token Answer, Five Ways

Same task for every architecture: read 4 context tokens, emit a 4-token answer. Watch where the state lives at each step.

THE TASK ctx₁ ctx₂ ctx₃ ctx₄ → read all four → a₁ a₂ a₃ a₄ ① SOFTMAX ATTENTION — full transcript replay aₜ needs scores vs EVERY past token: state = KV cache, grows O(s) a₁ KV CACHE — stores K,V for all 5 tokens replayable transcript · grows forever ② SSM (Mamba) — rolling notebook hₜ = Ā hₜ₋₁ + B̄ xₜ, Ā = exp(Δ·A) — one fixed-size h, updated per token h₀=0 →ᾲ h₁ →ᾲ h₂ →ᾲ h₃ →ᾲ h₄ → a₁…a₄ STATE = one h, size d — never grows ③ LINEAR ATTENTION (Gated DeltaNet) — matrix notebook Sₜ = Sₜ₋₁(I − β k kᵀ) + β v kᵀ — erase-then-write into a d×d state matrix S : d × d matrix ←k,v updates each token STATE = fixed d×d — can overwrite, can't append ④ DIFFUSION LM — carve all 4 answers from noise, in parallel ?₁ ?₂ ?₃ ?₄ ⇒ refine all 4 jointly, N denoise rounds ⇒ STATE = the whole noisy sequence ⑤ DECISION MODEL (JEV) — encode ctx ONCE, read probabilities off hidden states Yes 80% / No 20% shared-context prefill → parallel masked question branches → typed readout, no decode loop

The visual tells the whole story: attention replays the transcript (state grows), SSM and linear attention keep a fixed notebook (state compresses), diffusion works on the whole answer at once, and the decision model never generates at all. Compression is the theme — and every compression is a bet about what the task won't need later.

05 The Time-Complexity Table

For s sequence length, w window size, d state dimension, n generated tokens:

RegimeTrainingDecode stepTotal for n tokensWhere state lives
Full attentionO(s²·d)O(s·d) — read whole KVO(n·s·d)KV cache (grows)
Sliding windowO(s·w·d)O(w·d)O(n·w·d)window KV (capped)
SSM / linear attentionO(s·d²)O(d²) — one state updateO(n·d²)fixed h or S
Decision readout (JEV)one O(S) prefill~O(1) per question branch~O(S) for Q questionsshared KV, reused

Read the last row against row one: Q questions over S context tokens cost ~Q·S in decode-heavy form versus ~S with a single shared prefill. That asymmetry — encode once, branch many — is the entire JEV pitch, and it's the same asymmetry that makes prefix caching cheap for chat.

06 The Families, In Turn

6.1 State-Space Models — Mamba

A linear recurrence that replaces attention's global lookup with a compressed rolling state. The selective trick: the transition matrix depends on the input, so the model can choose per-token what to remember and what to forget.

hₜ = Ā hₜ₋₁ + B̄ xₜ yₜ = C hₜ Ā = exp(Δ·A), B̄, C, Δ = f(xₜ) ← input-dependent (selective) scan cost: O(s) vs attention's O(s²)

Δ is the gate. Large Δ → Ā ≈ 0 → previous state erased, current token written sharply (a "reset" on a new section or a fact worth anchoring). Small Δ → state persists (drift over filler). Compression becomes inference: deciding what h keeps is a learned judgment call about the future, made at read time.

Mamba-2's SSD duality shows the recurrence can be rewritten as a form of (masked, decayed) linear attention — the two families are two views of one algebra. That matters because it means hybrid layers (next subsection) are mixing siblings, not strangers.

Why it may win

Linear time, constant memory per layer, real length extrapolation, throughput that crushes attention on long sequences. On language modeling perplexity it matches transformers at matched scale.

Why it won't (yet)

The multi-query recall anchor problem: tasks like MQAR — "what were the exact values associated with keys K₁, K₃, K₇?" — require retrieving specific past tokens, and a fixed-size h has thrown away exactly that precision. Compression destroys addressable memory.

6.2 Attention–SSM Hybrids — Jamba, Zamba, Griffin/Hawk

The engineering compromise: interleave. A few full-attention layers (exact recall, exact copy) among many SSM/linear layers (cheap drift), so recall anchors exist without paying O(s²) everywhere.

layers: [SSM] [SSM] [ATTN] [SSM] [SSM] [ATTN] … (e.g. 1:7 or 1:5 ratio) Jamba: attention + Mamba + MoE in one block → 1 attention layer per ~8

Why hybrids exist — the per-layer tradeoff: each layer independently chooses between compute/memory spend (attention: O(s) state, exact recall) and bandwidth thrift (SSM: fixed state, lossy recall). Recall behavior is task-level, but the cost is layer-level — so spend attention only where the representation actually needs addressable memory. Jamba (AI21), Zamba (Zyphra), and Griffin/Hawk (DeepMind) all landed near-attention quality at SSM-class cost on long context.

Verdict — the most deployment-realistic challenger; if long-context models drift this way, you'll mostly notice as cheaper KV bills.

6.3 RWKV v5/v6

A linear-attention RNN wearing an LSTM's coat, famous for community training at real scale. v5/v6 make the recurrence data-dependent (like Mamba's Δ):

token = (R)eceptance, (W)eight, (K)ey, (V)alue state: wₜ = μₜ ⊙ wₜ₋₁ + kₜ (per-channel decay, input-gated) output: yₜ = rₜ ⊙ Σ (w ⊙ kᵢ) vᵢ over the state

The bytes-at-runtime framing: a deployed RWKV runs as a tiny stateful kernel — weights + one small state vector per layer. No KV cache file, no context-length allocator. That makes it the natural architecture for on-device and edge inference, where the KV cache is the memory problem (doc 22).

Verdict — edge-first bet; same recall ceiling as all fixed-state models.

6.4 Gated DeltaNet & the Linear-Attention Family

Linear attention (Katharopoulos et al.) replaced softmax(q·kᵀ)·v with φ(q)·(Σ φ(k)ᵀ v) — associative, O(s), but a sum that only grows. DeltaNet adds a delta rule: retrieve-then-correct the state, so newer keys can overwrite stale ones instead of just adding:

Sₜ = Sₜ₋₁(I − βₜ kₜ kₜᵀ) + βₜ vₜ kₜᵀ └─ erase: remove old (kₜ → ·) value ─┘ └ write: vₜ at key kₜ ┘ βₜ ∈ (0,2) — a learned gate between pure accumulation (β→0) and hard overwrite (β→1)

The MQAR lesson: associative recall (multi-query associative recall) separates from perplexity. A linear-attention model can match a transformer's next-token loss while failing MQAR, because perplexity is dominated by local, drift-y prediction where a compressed state suffices — while MQAR demands exact key→value lookup that compression can't preserve. Benchmark on recall, not just perplexity, when evaluating any fixed-state model.

Verdict — the delta rule closed much of the recall gap; hybrids still hold the edge.

6.5 xLSTM — Exponential Gating, LSTM Form

Beck et al. (2024) asked: what if the LSTM's problem was never gating, just capacity and parallelism? xLSTM recovers exponential memory decay (like Mamba's Ā = exp(Δ·A)) inside LSTM form: sLSTM adds scalar exponential gates with new mixing, mLSTM swaps the scalar cell for a matrix memory cell with a covariance-style update — effectively a linear-attention state with LSTM-style gating, fully parallelizable.

cₜ = fₜ ⊙ cₜ₋₁ + iₜ ⊙ vₜkₜᵀ (matrix cell, fₜ = exp-gated forget) nₜ = fₜ ⊙ nₜ₋₁ + iₜ kₜ (normalizer state)
Verdict — proof that the pre-transformer lineage, upgraded, is competitive at small–mid scale; frontier traction unproven.

6.6 BitNet b1.58 — The Ternary Bet

Not a sequence-model change — an arithmetic change to the whole stack. Every weight is forced to −1, 0, or +1; activations stay INT8:

Q(x) = clip( round( x / s ), −1, 0, 1 ) s = mean(|x|) scale info/weight: log₂ 3 ≈ 1.58 bits → "b1.58"

Why it matters for a hardware-minded reader: a matmul over {−1,0,1} weights has no multiplies — it's adds and sign flips — so the FLOP cost per weight drops ~9× versus FP16, and so does the memory traffic that doc 10 identified as the decode bottleneck. Training needs tricks (this is the hard part): quantization-aware training from scratch, and spot precision — keeping critical accumulation steps in higher precision so gradients stay stable through the rounding.

Verdict — orthogonal to everything else (works on transformers, hybrids, anything); a bandwidth multiplier, not a paradigm. KV cache remains FP16 — for now.

6.7 Byte-Level Models — Byte Latent Transformer

BLT (Meta, 2024) deletes the tokenizer. Raw bytes flow in; a small local model groups them into patches whose size tracks local entropy; a big global transformer attends over patches; a local decoder expands back to bytes.

patch size p ∝ entropy E(x): "the cat sat" → big patches (predictable) "E=mc²" / "json_field_9" → tiny patches (surprising) p_count vs token count: compute scales with information, not BPE luck

The trade: compute per byte moves up front — you pay a small-model pass over every byte before the big model sees anything. But you gain: no OOV, no tokenization artifacts in code/math, no "why is my French prompt 30% more expensive in tokens" tax. At scale, BLT matched Llama-3 tokenizer quality at better byte-level compute parity.

Verdict — a fairness and robustness win; a latency cost paid before the expensive model even runs.

6.8 Diffusion LMs — LLaDA, Mercury/Grok-Fast Style

Autoregressive models factor text as a chain: p ∝ Πₜ p(xₜ | x<t) — one token at a time, forever. Diffusion LMs factor it as a denoising problem: corrupt the whole sequence with mask noise, learn to restore it, refine all positions in parallel:

forward (corrupt): q(xₜ | x₀) = N( √ᾱₜ · x₀, (1 − ᾱₜ)·I ) learned reverse: denoiser predicts x₀ from noisy xₜ; iterate until clean per step: MANY tokens refined vs exactly ONE emitted

This is the bridge to doc 24: speculative decoding and multi-token heads try to squeeze parallelism into an autoregressive loop; diffusion LMs are born parallel — several refinement rounds emit a whole block. Mercury (Inception Labs) reportedly runs 5–10× faster than size-matched autoregressive models precisely because each forward pass advances many tokens. The cost: refinement rounds re-run full attention over the whole sequence (wasted compute on already-finalized tokens), and quality on long, tightly-ordered text still trails the best autoregressive models.

Verdict — the speed story is real; the frontier-quality story is still open.

6.9 JEPA — Predict Latents, Not Tokens

The LeCun-lineage answer to a different question: what if generation is the wrong objective? Joint-Embedding Predictive Architectures (I-JEPA on images, V-JEPA on video) predict abstract representations of missing content, never pixels or tokens:

E = ‖ f(x_target) − D( g(x_context) ) ‖² in latent space no decoder, no pixel/token reconstruction, no generative loss

By predicting in representation space, the model is never forced to spend capacity modeling unpredictable detail — the exact noise that makes pixel/token generators hallucinate texture. Why it may matter for LLMs: "token detokenization" — the waste of modeling surface word choice rather than meaning — is plausibly the same failure; a JEPA-style LLM would score candidate thoughts, not spellings. Why it's not here: without a generative head you can't sample text from it — the interface itself has to change, which is a product problem, not just a research one.

6.10 JEV (TypeSafe) — The Decision Model ★

The newest and strangest entry — a typed-decision architecture. Instead of generating an explanation ending in "Yes", it outputs the decision itself with a calibrated probability: Yes 80% / No 20% — read directly from hidden states over the allowed answers. There is no autoregressive decoding loop at all.

Honest framing: this is a new, poorly documented design — an illness of an ecosystem where closed internals force black-box reconstruction from ~10K API calls. Everything below is a public-article reconstruction; treat "reportedly" as attached to every claim.

reported architecture (black-box reconstruction): 1. causal transformer encodes shared context ONCE → single prefill 2. per layer: KV/state of that prefill is REUSED across many question branches (attention masks block sibling branches — each branch sees shared context + its own option set only) 3. branches run in parallel, isolated 4. probabilities read out over the allowed answers — no decode loop shared-context math: Q questions over S context tokens decode-style: ~ Q·S tokens of state work JEV-style: ~ S (prefill once, reuse across all Q branches)

Calibration as the training target. Reportedly trained with RLCD (Reinforcement Learning for Calibrated Decisions) — plausibly a log-loss or Brier-style objective — so that a stated "80%" event fires about 80% of the time. Reported MMLU calibration error ~0.031, with most predictions concentrated at high confidence. A transformer's sampled text has no such guarantee: "I'm 80% sure" in prose is theater; a calibrated readout is a contract.

Expected-cost routing — why agents care. Given a decision with probability p of the bad branch, escalate-to-human (or to a bigger model) when the expected cost of trusting it exceeds the escape cost:

escalate iff p·C_miss > (1 − p)·C_escape example: C_escape = 1 (asking costs 1), C_miss = 9 (missed urgent = 9) → escalate when 9p > 1 − p → p > 0.1

With calibrated p, routing thresholds become arithmetic instead of vibes. A fleet of agent decisions — approve, retry, escalate — becomes a budget line you can actually compute. This connects directly to the doc-21 theme: spend expensive compute only where the expected cost justifies it.

Verdict — not a text-model competitor; a new interface category (typed, calibrated decisions) built on the same transformer backbone. Watch it: if calibration holds at scale, agent harnesses get a component they've lacked since day one.

6.11 Three Footnotes That Could Become Chapters

Multi-token heads

Doc 02's trick: parallel output heads predicting tokens t+1…t+k in one pass. Cheap 2–3× decode speedup; full survey of the strategy space in doc 24.

Memory layers

Product-key memory: a huge parameter store consulted by nearest-neighbor lookup — a few keys in, one value out. Adds billions of "lookup parameters" without quadratic attention; recent work shows it beating MoE at matched FLOPs for fact-heavy tasks.

Titan-style test-time memory

Memorize at inference: an online surprise gradient trains a small memory module while the model runs — loss Δ = Σ ‖ f(x) − g(h, x) ‖² updated per token. Learning at test time; caching at inference time (doc 07) as its learned cousin.

07 Will They Replace the Transformer?

Probably not by revolution. By infiltration.

The incumbent's moat is not mathematical — Mamba and DeltaNet are legitimately better sequence math in several regimes. The moat is infrastructural:

Silicon. GPUs and TPUs are tuned for dense matmuls — exactly attention's workload. SSM scans, byte patchers, and ternary kernels run on the same chips at a fraction of their potential until custom kernels (and possibly custom silicon) mature.
Tooling. Every serving stack — vLLM, TensorRT-LLM, batching, quantization, doc 10's whole optimization ladder — assumes a KV cache. A deployed Mamba has no KV cache to manage: great for memory, but ten thousand lines of serving logic don't apply.
The KV cache is infrastructure. Prefix caching, doc 22's cache-type zoo, and shared-context tricks like JEV's all monetize one artifact: the reusable KV state. An architecture without it must reinvent an equivalent reuse story before the economics close.
Trained workflows. Prompting, fine-tuning, eval harnesses, agent frameworks — all calibrated against autoregressive text. A diffusion or decision model changes the interface, and interfaces are where switching costs live.
🔮
The likely ending: the transformer survives as a name for a stack that is no longer purely transformer — attention layers at recall-critical points, SSM/linear layers everywhere else (Jamba-style), ternary or low-bit weights (BitNet-style), byte patching under the hood, diffusion or multi-token emission on the way out, and typed decision readouts beside the text interface. Each challenger contributes one organ. Doc 02's model doesn't die; it gets dissected.

08 Mental Models

Transcript vs notebook vs form

Transformer = full transcript replay — every word is still on the table, reread each step. SSM = rolling notebook — you keep one page of distilled notes and update it as you read; fast, but you can't quote the original verbatim. JEV = typed form over the transcript — read the document once, then tick pre-printed checkboxes with confidence scores; you never write prose at all.

Notebooks can't answer "quote the exact third sentence"; transcripts can. Forms can't explain themselves — the reasoning must be trained into the readout.
Compression-as-inference

Every fixed-state model (Mamba, RWKV, DeltaNet, xLSTM) is a learned compression algorithm running online: h is a zip of history, and the update rule decides what's worth keeping. Compression ratio is the whole ballgame — too lossy and recall breaks (MQAR), too faithful and you've reinvented the KV cache.

The model can't know at read time which details a future query will need — recall failures are inherent mis-predictions of future relevance, not bugs.
Entropy as a budget dial

BLT spends compute ∝ surprise; Mamba's Δ gates memory ∝ surprise; Titan memorizes ∝ surprise. Three architectures, one instinct: predictable bytes deserve cheap handling, surprising bytes deserve expensive handling — the same instinct as doc 07's cached-vs-fresh pricing.

Entropy is model-relative — one model's surprise is another's old news; the dial needs calibration per domain.

09 Common Misconceptions

"SSMs are just worse transformers." They're different points on the recall-vs-cost curve. For pure sequence drift they match or beat attention; they lose specifically on exact multi-query recall. "Worse" without naming the task is meaningless.

"Linear time means faster responses for users." Not automatically: decode speed is bandwidth-bound (doc 10), and a fixed-size state must be read and rewritten every step. The O(s) win shows up at long context and in training/throughput, not necessarily in time-to-first-token on a 2K prompt.

"BitNet means you can quantize your existing model to 1.58 bits." No — b1.58 models are trained from scratch with quantization-aware objectives. Post-training-rounding an FP16 model to three values destroys it.

"Diffusion LMs are non-autoregressive, so they dodge the quality penalty." They trade it: parallelism for refinement waste and weaker long-range ordering. Speed ≠ quality; Mercury's 5–10× is real, but so is the frontier gap on long-form coherence.

"JEPA/decision models can replace LLMs." They abandon the generative interface — no free-form text comes out. They're complements: representation scorers and decision readouts beside a generative core, not instead of one.

🌅
Closing insight: the transformer won 2017–2024 because it was the best compression of a transcript into parallelizable state. Every architecture in this doc accepts that framing and fights over the terms — how much to compress, where, in how many bits, and whether the output must be text at all. The next decade's stack will likely be a chimera of all of them: hybrid layers for the math, ternary weights for the bandwidth, bytes underneath, parallel emission above — and, if JEV's calibration survives scrutiny, a typed-form interface where agents need decisions rather than essays. Track the churn in News · Latest AI Enhancements.